Genome-Wide Data-Mining of Candidate Human Splice Translational Efficiency Polymorphisms (STEPs) and an Online Database
نویسندگان
چکیده
BACKGROUND Variation in pre-mRNA splicing is common and in some cases caused by genetic variants in intronic splicing motifs. Recent studies into the insulin gene (INS) discovered a polymorphism in a 5' non-coding intron that influences the likelihood of intron retention in the final mRNA, extending the 5' untranslated region and maintaining protein quality. Retention was also associated with increased insulin levels, suggesting that such variants--splice translational efficiency polymorphisms (STEPs)--may relate to disease phenotypes through differential protein expression. We set out to explore the prevalence of STEPs in the human genome and validate this new category of protein quantitative trait loci (pQTL) using publicly available data. METHODOLOGY/PRINCIPAL FINDINGS Gene transcript and variant data were collected and mined for candidate STEPs in motif regions. Sequences from transcripts containing potential STEPs were analysed for evidence of splice site recognition and an effect in expressed sequence tags (ESTs). 16 publicly released genome-wide association data sets of common diseases were searched for association to candidate polymorphisms with HapMap frequency data. Our study found 3324 candidate STEPs lying in motif sequences of 5' non-coding introns and further mining revealed 170 with transcript evidence of intron retention. 21 potential STEPs had EST evidence of intron retention or exon extension, as well as population frequency data for comparison. CONCLUSIONS/SIGNIFICANCE Results suggest that the insulin STEP was not a unique example and that many STEPs may occur genome-wide with potentially causal effects in complex disease. An online database of STEPs is freely accessible at http://dbstep.genes.org.uk/.
منابع مشابه
SPOT: a web-based tool for using biological databases to prioritize SNPs after a genome-wide association study
SPOT (http://spot.cgsmd.isi.edu), the SNP prioritization online tool, is a web site for integrating biological databases into the prioritization of single nucleotide polymorphisms (SNPs) for further study after a genome-wide association study (GWAS). Typically, the next step after a GWAS is to genotype the top signals in an independent replication sample. Investigators will often incorporate in...
متن کاملRNA Structures Affected By Single Nucleotide Polymorphisms In Transcribed Regions Of The Human Genome
Single-stranded RNAs fold into base-paired structures sometimes critical for RNA function.We report the first genome-wide computational analysis of mRNA structures, obtaining predicted minimum free energy structures (MFE) and ensembles of top-scoring structures.Evaluation of mRNA structures from >12,450 genes indicates that thermodynamically favorable structures preferentially form within speci...
متن کاملSNP mining porcine ESTs with MAVIANT, a novel tool for SNP evaluation and annotation
MOTIVATION Single nucleotide polymorphisms (SNPs) analysis is an important means to study genetic variation. A fast and cost-efficient approach to identify large numbers of novel candidates is the SNP mining of large scale sequencing projects. The increasing availability of sequence trace data in public repositories makes it feasible to evaluate SNP predictions on the DNA chromatogram level. MA...
متن کاملروشی کارا برای کاوش مجموعه اقلام پرتکرار در تحلیل دادههای سبد خرید
Discovery of hidden and valuable knowledge from large data warehouses is an important research area and has attracted the attention of many researchers in recent years. Most of Association Rule Mining (ARM) algorithms start by searching for frequent itemsets by scanning the whole database repeatedly and enumerating the occurrences of each candidate itemset. In data mining problems, the size of ...
متن کاملPolymiRTS Database 2.0: linking polymorphisms in microRNA target sites with human diseases and complex traits
The polymorphism in microRNA target site (PolymiRTS) database aims to identify single-nucleotide polymorphisms (SNPs) that affect miRNA targeting in human and mouse. These polymorphisms can disrupt the regulation of gene expression by miRNAs and are candidate genetic variants responsible for transcriptional and phenotypic variation. The database is therefore organized to provide links between S...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره 5 شماره
صفحات -
تاریخ انتشار 2010